Papers by Shei Pern Chua
Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) safety alignment predominantly operates on a binary assumption that requests are either safe or unsafe. |
| Approach: | They propose a methodology that embeds harmful requests within ethical framings to exploit this vulnerability. |
| Outcome: | The proposed framework achieves high success rates by exploiting model's own ethical reasoning to frame harmful actions as morally necessary compromises. |